Repository navigation
[Nemotron][CUDA] Add deterministic moe_router_dispatch - #504
Draft
RichApple123 wants to merge 1 commit into
Draft
RichApple123 wants to merge 1 commit into
RichApple123 wants to merge 1 commit into
Conversation
Signed-off-by: RichApple123 <134492660+RichApple123@users.noreply.github.com>
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true
Comment |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements the SM90 CUDA
moe_router_dispatchwork item of #434: FP32 projection, sigmoid, bias-corrected top-six selection, normalization and deterministic token packing, with backward for routing weights and packed payloads. Supports BF16/FP32 inputs with FP32 router weights at Nano's 2688-hidden, 128-expert geometry.Adds a numerical reference, CPU/GPU correctness and batch-invariance tests, a cuBLAS + Transformer Engine benchmark, documentation and CI wiring for
test-nemotron. The ordering and fixed arithmetic policy are proposed for section 4.4 review. Expert MLP/combine and distributed communication are outside this change.Validation:
The target branch still has the eight documentation link warnings addressed by #458 on main. This change adds no documentation warnings; strict docs passes locally with a separate fix for those links. Target-branch CI and maintainer approval of the proposed contract remain pending.